Papers with French dataset
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)
Copied to clipboard
Cécile Macaire, Chloé Dion, Jordan Arrigo, Claire Lemaire, Emmanuelle Esperança-Rodier, Benjamin Lecouteux, Didier Schwab
| Challenge: | Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments. |
| Approach: | They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions. |
| Outcome: | The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence. |
BLM-AgrF: A New French Benchmark to Investigate Generalization of Agreement in Neural Networks (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing benchmarks for deep learning are based on massive amounts of data, which are effective in hiding some of the shallowness of the learned models. |
| Approach: | They propose to use a French dataset to learn the underlying rules of subject-verb agreement in sentences, inspired by visual IQ tests known as Raven’s Progressive Matrices. |
| Outcome: | The proposed method is based on Raven’s Progressive Matrices, a visual IQ test, and a dataset built using the BLM framework. |
HISTOIRESMORALES: A French Dataset for Assessing Moral Alignment (2025.naacl-long)
Copied to clipboard
Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin, Charlotte Laclau, Guillaume Metzler, Christophe Gravier
| Challenge: | HistoiresMorales is a dataset based on moralStories in French . it is based upon annotations of moral values within the dataset . |
| Approach: | They propose a dataset in French that aims to align language models with moral values . they use annotations to ensure their alignment with French norms . |
| Outcome: | The proposed dataset guarantees grammatical accuracy and adaptation to the French cultural context. |
Multilingual prediction of Alzheimer’s disease through domain adaptation and concept-based language modelling (N19-1)
Copied to clipboard
Kathleen C. Fraser, Nicklas Linz, Bai Li, Kristina Lundholm Fors, Frank Rudzicz, Alexandra König, Jan Alexandersson, Philippe Robert, Dimitrios Kokkinakis
| Challenge: | Existing work on speech and language models has been limited by the size of available datasets. |
| Approach: | They propose to augment a small French dataset with a much larger English dataset to augment the language model to model the order in which information units are produced by dementia patients and controls. |
| Outcome: | The proposed model improves classification performance in English and French separately. |
He said “who’s gonna take care of your children when you are at ACL?”: Reported Sexist Acts are Not Sexist (2020.acl-main)
Copied to clipboard
Patricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari, Gloria Origgi, Marlène Coulomb-Gully
| Challenge: | Sexism is prejudice or discrimination based on a person's gender. |
| Approach: | They propose to use a French dataset annotated for sexism detection to characterize sexist content and to train deep learning experiments on tweets. |
| Outcome: | The proposed dataset is the first to be used for sexism detection in France and constitutes a first step towards offensive content moderation. |
Give me your Intentions, I’ll Predict our Actions: A Two-level Classification of Speech Acts for Crisis Management in Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | Using social networks, social media is a vital tool for emergency management and social media has been used to generate valuable information in crisis situations. |
| Approach: | They propose to measure for the first time the role of SA on urgency detection in tweets . they propose to use a two-layer annotation scheme to annotate tweets for both SA and urgency . |
| Outcome: | The proposed scheme combines two-layer annotation scheme and deep learning experiments to detect SA in a crisis corpus. |
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark (2024.lrec-main)
Copied to clipboard
| Challenge: | Intent classification and slot-filling tasks are essential tasks of Spoken Language Understanding (SLU). |
| Approach: | They propose to use a MEDIA SLU dataset to train a multilingual model to achieve both tasks jointly. |
| Outcome: | The proposed model can be trained on multiple datasets including the MEDIA dataset and extends to more tasks and use cases. |
CIS-BWE: Chaos-Informed Speech Bandwidth Extension (2026.acl-long)
Copied to clipboard
| Challenge: | CIS-BWE introduces two chaos-informed discriminators for capturing the deterministic chaos from speech. |
| Approach: | They propose a novel adversarial Bandwidth Extension framework that introduces two chaos-informed discriminators for capturing the deterministic chaos from speech. |
| Outcome: | The proposed framework achieves better performance across nine subjective and objective evaluation metrics with a 40x reduction in discriminator size and overall 0.5x fewer parameters, establishing a new baseline in the BWE task. |